Papers with speech encoder training
R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for speaker and noise-invariant speech representations use unlabeled audio data to pretrain encoders, generating good representations for downstream tasks like automatic speech recognition (ASR) and speaker identification. |
| Approach: | They propose a domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-in-variant clustering. |
| Outcome: | The proposed method reduces computational resources by 12X compared to state-of-the-art methods while outperforming them in severely distorted speech scenarios. |